Back

Molecular Systems Biology

Springer Science and Business Media LLC

Preprints posted in the last 30 days, ranked by how well they match Molecular Systems Biology's content profile, based on 162 papers previously published here. The average preprint has a 0.12% match score for this journal, so anything above that is already an above-average fit.

1
A generalized growth law for translation- and transcription-targeting antibiotics captures drug interactions

Gadjisade, N.; Mori, M.; Bollenbach, T.

2026-08-24 systems biology 10.64898/2026.08.21.746212 medRxiv
Top 0.1%
22.7%
Show abstract

Bacterial growth laws quantitatively connect intracellular resource allocation to growth rate, enabling accurate predictions of physiology and antibiotic responses. Yet these laws have been rigorously tested for only a handful of perturbations. Here, we show that the growth law linking ribosome levels to growth rate under translation-inhibiting antibiotics is not universal, but rather depends on the antibiotic's mechanism of action. Quantitative proteomics across finely resolved one- and two-dimensional antibiotic gradients showed that inhibitors of translocation elongation or peptide bond formation elicit the canonical rise in ribosome levels, consistent with the growth law. By contrast, antibiotics disrupting translation initiation or fidelity produced distinct responses without ribosome upregulation. The transcription inhibitor rifampicin even reduced ribosome abundance. Combining antibiotics with divergent ribosome responses revealed a generalized growth law, in which the individual responses to perturbations superimpose. Embedding this law in a mathematical model explains distinct drug interaction patterns observed between rifampicin and different translation inhibitors. A low-dimensional structure pervades the entire proteome, enabling prediction of responses to drug pairs based on single-drug measurements. Together, these findings broaden the scope of bacterial growth laws and provide new principles for predicting responses to antibiotic combinations.

2
Biphasic Temporal Remodeling Of The Proteome In A Polyglutamine-Expanded Huntingtin In Vitro Aggregation Cell Model: From Early Rna-Regulatory Compensation To Selective Mitochondrial Energy Failure

Sonmez, E.; Mutlu, P.; Ozlevent, C.; Sarihan, M.; Akpinar, G.; Kasap, M.; Cimen, H.

2026-08-24 systems biology 10.64898/2026.08.21.746221 medRxiv
Top 0.1%
15.3%
Show abstract

Huntington disease (HD) is caused by a polyglutamine expanded huntingtin protein that exerts progressive cellular toxicity. However, the temporal sequence of pathogenic, particularly early and reversible versus late and irreversible events remain incompletely defined, despite their distinct therapeutic implications. To delineate this trajectory, we profiled the proteome of a huntingtin expressing cell model at early (72 h) and late (144 h) stages. Rather than a linear progression, pathogenicity unfolded in two discrete phases. At the early stage, cells exhibited a broad activation of RNA processing, splicing, and protein synthesis machinery, consistent with an adaptive response aimed at preserving gene expression fidelity under stress. By the late stage, this compensatory program had collapsed, giving rise to a dominant failure in mitochondrial energy metabolism. Notably, 85% of proteins altered at both time points reversed direction of change between stages, indicating that mutant huntingtin reprograms cellular function wholesale rather than amplifying a fixed set of perturbations. Detailed analysis of mitochondrial respiratory complexes revealed that terminal ATP generating components (cytochrome c oxidase and ATP synthase) were severely affected, whereas upstream electron transport elements were retained or upregulated. Leveraging this proteomic map, we applied an AI assisted, direction aware drug repurposing strategy. Of 1,712 differentially expressed proteins, 498 were druggable, and 89 mapped to approved agents with mechanisms concordant with the required correction. These included Complex I targeted agents (metformin, ME 344) and mitochondria directed therapeutics (SS 31, MitoQ), several of which have previously been evaluated in HD. Collectively, these findings define a biphasic course of huntingtin toxicity and highlight an early therapeutic window in which intervention is most likely to be applied, prior to irreversible deterioration of mitochondrial respiratory function.

3
Dissecting context-dependent cancer vulnerabilities using Perturb-seq

Maffa, S.; Boyle, I. A.; Ward, L.; Colgan, W. N.; Borck, P.; Simerzin, A.; Adeagbo, A.; Olajide, O.; Wie, S.; Liang, H.; Wienand, K.; Shibue, T.; Ray, J.; Paolella, B.; Campbell, C. D.; Vazquez, F.; Dempster, J. M.

2026-08-25 cancer biology 10.64898/2026.08.24.746802 medRxiv
Top 0.1%
12.7%
Show abstract

Background CRISPR-mediated viability assays in diverse cancer cell lines have informed cancer biology and precision medicine, but cell fitness is not the only cancer-relevant phenotype. Gene expression profiling provides insight into cellular stress, inflammation, and differential state, while still identifying activation of cell-death pathways. Perturb-seq allows scalable functional genomics screening of expression phenotypes at single-cell resolution, however existing datasets cover only a small number of work-horse cell lines. Results We produced a proof-of-concept Perturb-seq dataset targeting 100 genes in 16 diverse cancer cell lines. In the process, we established methods to address single-cell technical artifacts, identified Cas9-mediated chromosomal aberrations and assessed screen quality. Even with a limited library, we observed common signatures of deleting essential genes as well as context-specific responses based on intrinsic genomic properties of the models. For example, we inferred a previously undescribed relationship between dependence on the ER-golgi transport gene immediate early response 3 interacting protein 1 (IER3IP1) and oxidative stress, demonstrating the potential of integrated Perturb-seq for hypothesis generation. Conclusions We established a framework for building a comprehensive map of post-perturbational transcriptional phenotypes using parallel Perturb-seq experiments across multiple cell lines. We demonstrated that integrated Perturb-seq experiments spanning diverse contexts enable hypotheses about gene function specific to tissue types or cancer subtypes - suggesting large-scale, genome-wide datasets would offer invaluable insight into the highly context-dependent nature of cancer biology.

4
Audited vibe coding suggests partial fetal-like convergence of tumor proteomes

Meyer, J. G.

2026-08-31 cancer biology 10.64898/2026.08.26.745609 medRxiv
Top 0.1%
12.6%
Show abstract

The balance between how much human tumors recapitulate fetal tissue programs versus lose adult tissue identity remains unresolved. I used audited vibe coding, a human-mediated, cross-model critique-and-refinement workflow, to re-analyze a public pan-cancer proteomic atlas. A primary large language model wrote and executed the analysis under scientific direction, while a separate model family audited the code, outputs and claims; findings were returned for correction across seven versioned releases. Among 229 tumor-adjacent pairs in seven organs, tumor-minus-adjacent proteomic change partially aligned with reverse fetal-to-adult maturation (organ-balanced cosine, 0.240; 95% interval, 0.138 to 0.335), with positive alignment in 189 of 229 patients (82.5%). The organ-balanced projection coefficient was 0.195 (95% interval, 0.069 to 0.244), indicating movement along only part of the developmental distance. Although reverse maturation overlapped adult-identity loss, a positive developmental component remained after identity loss entered first (0.203; 95% interval, 0.129 to 0.239). Suppression of adult-high proteins contributed to more positive alignment than reactivation of fetal-high proteins. The vibe coding audits identified substantive defects. A common-mask correction reduced the matched-organ advantage from 0.074 to 0.059; a missing-value correction barely changed aggregate geometry but replaced 5 of the top 40 liver contributors; and coupled resampling repaired uncertainty accounting without changing patient scores. As with any single report, the "vibe reanalysis" biological results are candidate discoveries pending independent replication. The workflow is a single feasibility case, not a reliability benchmark, and shows how conversationally generated analysis can be made more inspectable when model-written code is treated as untrusted, versioned and subject to separate-model critique and executable checks.

5
The unexplainable plasma protein measurements from Proximity Extension Assays

Gräf, J. F.; Kurgan, N.; Rasmussen, S.

2026-08-11 genetic and genomic medicine 10.64898/2026.08.10.26360077 medRxiv
Top 0.1%
10.6%
Show abstract

Plasma proteomics assays aim to capture biologically interpretable circulating protein signals. Here we systematically assess Olink Explore targets with exceptionally low variance explainability by integrating genetic associations, cross-platform concordance and tissue and peptide atlases. We identify a subset of protein targets likely lacking robust plasma signals, highlighting assay- and tissue-specific limitations with implications for panel design, statistical power and interpretation of large-scale proteomics studies.

6
Decoding Tumour-Specific Rewiring and Synthetic Lethality Through Genome-Scale Metabolic Models

Ibrahim, M.; Bhoite, R.; Lakshmanan, M.; Raman, K.

2026-08-21 systems biology 10.64898/2026.08.21.746174 medRxiv
Top 0.1%
10.6%
Show abstract

Cancer cells rapidly rewire their metabolism, from efficient energy production toward anabolic processes, to sustain uncontrolled growth. Decoding such metabolic shifts is essential for uncovering novel therapeutic targets. To map systems-level metabolic changes across cancer types, we built context-specific genome-scale metabolic models for eight tissues (lung, thyroid, stomach, prostate, liver, kidney, colon, and breast) using gene expression data from The Cancer Genome Atlas (TCGA). Applying constraint-based modelling, we then identified differentially regulated pathways through flux enrichment analysis, revealing tissue-specific rewiring: branched chain amino acid metabolism was suppressed in breast cancer; sphingolipid metabolism was downregulated in colon, kidney, and thyroid but upregulated in breast. We further propose a model-driven pipeline to identify and characterise metabolic vulnerabilities. We first identify synthetic lethal reactions in normal tissues and their corresponding single lethal counterparts in cancers, thereby enabling the identification of metabolic "collateral lethal" reaction pairs for each cancer. Model-predicted collateral lethal gene pairs, including CMPK1-AK in colon, ALDOA-PGD in prostate, and SLC25A26-UQCRB in liver models, were supported through computational validation using DepMap data on gene essentiality. Subsequently, we show how to interpret metabolic rewiring in cancer tissues while accounting for any collateral lethal pairs. In summary, our results establish a systemic framework for decoding metabolic rewiring and synthetic lethal vulnerabilities in cancer.

7
A confound-diagnostic toolkit for in silico perturbation with single-cell foundation models

Qiu, R.; Zhao, M. M.

2026-08-07 bioinformatics 10.64898/2026.08.04.732812 medRxiv
Top 0.1%
9.9%
Show abstract

Deleting a gene token from a cells input sequence offers a convenient native strategy for in silico perturbation, but the resulting embedding delta may not represent a biological knockout response. Apparent effects can instead reflect gene identity, universal responsiveness, limited tokenization coverage, library-size contamination, or circular state scoring. Here, we present a confound-diagnostic framework combining held-out increment testing, responsiveness adjustment, coverage gating, library-size diagnostics, and de-circularized state-shift analysis, together with a numerically matched reimplementation of frozen Geneformers perturbation engine. Across Frangieh and Replogle datasets and linear and nonlinear readouts, the native embedding delta provided no reproducible held-out improvement beyond gene identity. Signal-injection calibration showed that the test detected injected residual signal, whereas native increments remained below its detection floor. Matched controls traced apparent positives to raw-count library-size structure, broad responsiveness, and self-referential scoring, while coverage constrained perturbation applicability and estimate stability without establishing biological specificity. This model-adaptable framework helps determine when foundation-model perturbation readouts warrant biological interpretation. MotivationFoundation-model in silico perturbation could predict perturbation effects when matched experimental data are unavailable. However, in zero-shot settings, embedding-derived responses may reflect gene identity, universal responsiveness, tokenization limits, library-size artifacts, or circular state scoring rather than biological knockout effects. We therefore developed a reusable confound-diagnostic framework that applies matched controls to test whether native perturbation readouts contain information beyond these confounds and warrant biological interpretation.

8
A discrete protein subset drives structure prediction discordance in orphan proteins

Eicholt, L. A.; Middendorf, L.

2026-08-25 bioinformatics 10.64898/2026.08.24.746756 medRxiv
Top 0.1%
9.6%
Show abstract

Structure and disorder predictors are increasingly used as decision-grade tools in protein engineering and in the analysis of newly emerged proteins, yet how the current state-of-the-art behaves on sequences outside the well-charted evolutionary space remains poorly characterised. We previously reported that AlphaFold2 confidence and the disorder predictor flDPnn produced discordant predictions for naturally evolved de novo Drosophila proteins and for shuffled sequences. Here, we revisit the comparison with AlphaFold3 and the best-performing disorder predictor PUNCH2 on the same sequence sets together with conserved Drosophila proteins and intrinsically disordered proteins. The discordance persists: pLDDT correlates positively with PUNCH2 disorder in random and de novo proteins and negatively with {beta}-strand fraction, opposite to the conserved and disordered baselines. A class-specific, score-defined driver subset jointly captures the unusual high-pLDDT, high-disorder, low-strand combination and contains 24.5% of de novo, 29.4% of random, 5.1% of conserved, and 1.3% of disordered proteins. Removing this subset normalises the correlations. A held-out classifier trained on architectural and compositional features that were not used in the driver definition recovers the subset, with helix and coil fraction, sequence length, entropy and hydropathy as the strongest predictors. The discordance is therefore not a sequence-class artefact but a localised, compositionally identifiable phenotype that current predictors handle in a non-canonical way - a concrete failure mode that protein designers and others working on sequences remote in sequence space should be aware of when relying on predictor outputs.

9
PerturbLDM: conditional latent diffusion for modelling single-cell perturbation responses

Yu, L.; Hsieh, K.-L.; Chu, Y.; Lan, Q.; Zhao, X.; Hsu, Y.-C.; Wood, C. S.; Rasmy, L.; Pilie, P. G.; Zhi, D.; Zhao, Z.; Jiang, X.; Dai, Y.

2026-08-12 bioinformatics 10.64898/2026.08.07.743610 medRxiv
Top 0.1%
9.1%
Show abstract

Single-cell perturbation profiling maps intervention-induced phenotypes, yet experiments measure only a fraction of the perturbation-context space. Learning context-dependent perturbation effects could enable response prediction beyond measured conditions. Here we introduce PerturbLDM, a latent-diffusion framework for conditional generation of single-cell transcriptional responses. Following Tahoe-100M pretraining, it predicted 13,942 held-out combinations of observed drugs, doses and cell lines more accurately than existing methods, with higher matched-control effect correlation than an additive marginal baseline in 95.2% of conditions. The Tahoe-100M-pretrained model was further used to rank PANACEA compounds by pathway similarity, placing shared-mechanism pairs among nearest neighbours. In smaller datasets, PerturbLDM generated a mid-gestational fetal-colon state with 67% lower gene-wise error than Squidiff, retaining the balance between absorptive and BEST4/OTOP2-like epithelial programmes. In PBMCs, it captured six of seven interferon and antiviral programmes and the interferon-associated FAO-OXPHOS programme more accurately than scGen. Together, these results support conditional response generation across data scales and biological settings.

10
Symbolic regression enables coarse-grained model discovery of intracellular signalling dynamics

de Pomereu, T.; Fröhlich, F.

2026-08-21 systems biology 10.64898/2026.08.20.745973 medRxiv
Top 0.1%
8.2%
Show abstract

Cells respond to their environment through protein networks often dysregulated in cancer, making dynamical modelling crucial. Limitations in experimental data and computational resources motivate coarse-graining methods to build low-dimensional descriptions. Yet classical approaches to coarse-grained modelling rely on strong assumptions, leaving it unclear when partial experimental observations support reduced descriptions of system dynamics. Here we show that symbolic regression (SR) provides a data-driven way to test whether, and how compactly, the dynamics of a signalling system coarse-grain over the measured variables, and, when they do, infers mechanistically interpretable models. In synthetic enzyme systems, SR recovers Michaelis-Menten kinetics for the two-step mechanism and under three-step extensions. As data quality is degraded, SR simplifies toward effective kinetic laws while preserving correct theoretical limits. Applied to published time-resolved ERK phosphorylation data, SR identifies compact phospho-ERK rate laws in selected cancer-relevant gene overexpression contexts, yielding interpretable kinetic effects. A sparse neural ODE baseline requires few inputs where SR succeeds, but on average more where it fails, indicating that, where a reduced model is learnable at all, SR failure is associated with more complex dynamics that a simple mathematical model cannot describe. Together, these findings establish symbolic regression as a way to test when a compact coarse-grained description is warranted, generating hypotheses where one holds and motivating potential new measurements where it does not.

11
scROMA: batch-aware pathway-activity inference and a ground-truth simulation framework for single-cell transcriptomics

Zhubanchaliyev, A.; Najm, M.; Laigle, V.; Bonnet, E.; Martignetti, L.

2026-08-10 systems biology 10.64898/2026.08.07.743516 medRxiv
Top 0.1%
7.8%
Show abstract

BackgroundPathway-activity analysis summarizes gene-level single-cell measurements into interpretable functional modules, but widely used methods lack an integrated significance framework, do not account for the batch effects that pervade multi-sample studies, and are not natively interoperable with Python-based workflows. The field also lacks simulation resources with ground-truth pathway activity for quantitative benchmarking. ResultsWe present scROMA, a singular-value-decomposition-based method that quantifies pathway activity as coordinated variation, with per-cell scores, per-gene contributions, and permutation-based significance, natively integrated with the Scanpy/AnnData ecosystem. Its batch-aware extension is, to our knowledge, the first to correct batch effects within the gene-set subspace rather than across the full transcriptome, isolating technical variation at the pathway level while preserving signal in other genes. We also release a generative simulation framework producing synthetic data with fully specified ground-truth activities. On simulated benchmarks scROMA is competitive across tasks, and under batch effects its batch-aware mode recovers ordinal pathway structure that full-transcriptome integration misses. Across cystic fibrosis airway, intestinal-organoid, breast cancer, and lung cancer datasets it recovers established biology while separating it from technical and inter-donor variation; in the intestinal-organoid atlas it reproducibly recovers an inflammatory program across donors, separates its sustained from transient components, and resolves cell-type-specific niche-factor targets. ConclusionsscROMA is open-source and released with the simulation framework and pre-generated benchmark datasets as a community resource, providing a scalable, statistically grounded, and batch-aware approach to pathway-level analysis in single-cell transcriptomics.

12
Point-in-time evidence and cross-area clinical precedent anticipate clinical entry across 100 focal areas: retrospective validation of the Intangia triage layer

Elliott, T. O.; Molnar, S.; Peeters, G.; Collart, O.

2026-08-24 bioinformatics 10.64898/2026.08.19.745722 medRxiv
Top 0.2%
7.8%
Show abstract

Early-opportunity teams face a combinatorial problem: once a focal target, mechanism or indication is fixed, the space of plausible partners runs to thousands of candidates per area. Intangia's triage layer ranks that space from point-in-time evidence (how much literature, patent and clinical activity a candidate pairing has accumulated, and whether the partner already has clinical precedent in other contexts) so that review starts where clinical activity is most likely to begin next. This preprint validates that capability retrospectively across 100 focal areas spanning drug targets, mechanisms and disease indications, replaying 24.1 million historically scored combination-years with every area scored by a model trained on the other 99 and never on itself. The headline is operational. At a twenty-partner review shortlist per focal area, the median area's four-year first-alert precision is 0.234, against a matched random-ranker median of 0.008: roughly one in four shortlisted partners subsequently entered the focal clinical context within four years, about 38 times each area's own background rate (95% CI 31 to 45). A panel-level permutation puts the result at p = 0.0005. Discrimination generalises: the full 13-feature specification reaches a median leave-one-focal-out ROC-AUC of 0.922 (95% CI 0.911 to 0.929), with no area below chance and all 100 areas beating their strongest count-based baseline. Shortlisted entrants are anticipated with a median observed lead of two years within the evaluation window, and three years (interquartile range one to five) once the window cap is removed and every realised entrant is counted. The core ranking is carried by two interpretable signal families: cumulative co-occurrence counts and leave-one-area-out clinical precedent. Burst detection serves a complementary role: it supplies the time-stamped, source-specific momentum evidence attached to every recommendation (what is accelerating, and why now) rather than additional ranking power. A conditional view of the same landscape ranks candidates with no cross-area precedent against one another, enriched relative to matched random ranking, supporting a lower-yield emerging-opportunities capability. Two worked examples, PD-1 combination immunotherapy and CTLA-4, are point-in-time historical replays of the same architecture in familiar territory, showing what an alert looked like with the dated evidence behind it. The endpoint throughout is first clinical entry, not clinical success; prospective validation is the next stage.

13
SpY-C: Supervised Learning of Phosphopeptide Sequence Constraints Enables Global Prediction of SH2 Domain Binding

Kandoor, A.; Silva Oliveira, A. C.; Machida, K.; Blagoev, B.; Naegle, K. M.

2026-08-21 systems biology 10.64898/2026.08.18.744962 medRxiv
Top 0.2%
7.7%
Show abstract

Tyrosine kinase signaling for cell development and homeostasis in multicelluar organisms and a major biochemical contribution is by driving interactions between phosphorylated tyrosines (pY) and SH2 domain containing proteins. This assembly is so important to driving cell outcomes that a wide variety of experimental and computational approaches have been used to understand which SH2-pY interactions occur, which still remains a challenge given the immensity (more than 45,000 pY and 120 SH2 domains in the human proteome). Based on biophysical constraints suggested by comprehensive contact mapping, here, we ask whether an approach might consider first asking if pY sequences conform to the shared rules of SH2 domain recognition by developing a classification approach that combines diverse training data. A wide range of validation suggests this approach, SpY-C, can classify pY sites as having the potential, or not, to be involved in SH2 domain interactions. We find that a relatively small set of representative SH2 binders, integrated from different experimental techniques, provides good classification. We use this classifier to annotate the human phosphoproteome and individual experiments, to explore the consequences of using super-SH2 domain reagents for pY enrichment, and to analyze the effects of mutations in altering pY site function. SpY-C provides a helpful step to more rapidly annotating pY function and for possibly improving machine learning approaches focused on specific SH2-pY interactions downstream of a first pass classification approach.

14
VariantFlux: A genotype-first modelling workflow for predicting the impact of genetic variations on human metabolism

Nazem-Bokaee, H.

2026-08-09 systems biology 10.64898/2026.08.03.740213 medRxiv
Top 0.2%
7.6%
Show abstract

Human genetic variation is a major determinant of organ metabolism, yet how naturally occurring variants shape quantitative metabolic phenotypes remains unclear. We present VariantFlux, a workflow that integrates ancestry-aware variant interpretation into genome-scale metabolic modelling to generate personalised, variant-constrained kidney reconstructions. Using the Human1 v1.19 model, we built a kidney-specific baseline model constrained by 482 metabolites and analysed 2,547 individuals from the 1000 Genomes Project, in whom [~]50% of metabolic genes were predicted damaging by at least three computational tools. These variants, whose burden differed subtly across ancestries, were translated into gene-dosage-anchored flux constraints for homozygous knockouts and graded heterozygous knockdowns. Despite widespread perturbation, >97% of models preserved baseline growth, indicating strong metabolic robustness. Yet individual genomes exhibited distinct flux-rewiring patterns, with frequent individual-specific gain-of-flux events and fewer shared loss-of-flux reactions. Limited ancestry clustering suggests metabolic responses are driven mainly by unique variant combinations. VariantFlux links human genomes to organ-level flux phenotypes, enabling precision medicine, pharmacogenomics, and disease risk prediction. Conceptual advanceWe present VariantFlux, a genotype-first framework that integrates predicted variant effects directly into genome-scale metabolic networks to generate personalised, organ-specific flux phenotypes. Unlike association-based metabolomics studies, this bottom-up approach enables exploratory, mechanistic prediction of how naturally occurring genetic variation reshapes human metabolism.

15
Structural proteomics reveals a coagulation-complement accessibility signature of macrovascular invasion in hepatocellular carcinoma

Son, A.; Hur, M. H.; Cho, E. J.; Ji, J.; Han, E.; Choi, Y.; Park, J.; Lee, H.; Park, S.; Yu, S. J.; Kim, H.

2026-08-13 systems biology 10.64898/2026.08.12.744566 medRxiv
Top 0.2%
6.7%
Show abstract

Macrovascular invasion (MVI) and extrahepatic spread (EHS) define the most aggressive, treatment-refractory hepatocellular carcinoma (HCC), yet blood-based markers that report the underlying protein-network biology are lacking. Conventional proteomics measures protein abundance but not the conformational and protein-protein-interaction (PPI) states that govern function. We applied covalent proteome painting (CPP)--a dimethylation-based accessibility assay that reads out binding-site openness--to matched tumor and serum, reasoning that intravascular tumor dissemination remodels plasma protein complexes in a manner detectable as changes in accessibility. Eight treatment-native HCC patients were profiled by CPP using matched FFPE tumor and top-14- depleted serum on a Q Exactive Orbitrap HF. The 85 tumor-serum common proteins defined an 81-protein targeted panel, validated by multiple-reaction-monitoring (MRM) mass spectrometry with heavy stable-isotope-standard peptides (296 peptides; 3,717 light/heavy transition pairs) in 22 FFPE tumors and 22 matched sera. Accessibility was the light/heavy ratio (high, open; low, closed). We assessed differential accessibility, serum-tissue translatability, pathway enrichment, and biomarker/survival performance. Aggressive disease showed broadly decreased protein accessibility. MVI-associated changes were directionally concordant between tumor and serum (Spearman {rho}=0.21; 59% concordant), driven by coagulation and complement proteins (FGG, CTSD, LBP, C4BPA); the EHS axis did not translate. Decreased-accessibility proteins were enriched for complement-coagulation cascades and IGF/IGFBP transport. A six-protein serum accessibility signature discriminated MVI (leave-one-out cross-validated AUC 0.80; best single markers ceruloplasmin 0.83 and haemoglobin- 0.77), and MVI status trended with shorter overall survival (log-rank p=0.06). Accessibility-based serum proteomics captures MVI-associated protein-complex remodeling that abundance assays miss, nominating a coagulation/complement-anchored serum signature for vascular-invasive HCC that warrants prospective validation.

16
Cytokine interaction networks, not individual cytokines, drive anti-TNF response in Crohn's disease

Olbei, M.; Thomas, J. P.; Liu, Y.; Malas, S.; Modos, D.; Powell, N.; Korcsmaros, T.

2026-08-11 systems biology 10.64898/2026.08.11.744137 medRxiv
Top 0.2%
6.6%
Show abstract

Crohns disease (CD) is a chronic inflammatory condition of the gastrointestinal tract for which anti-tumour necrosis factor (anti-TNF) agents remain a first-line biologic therapy. However, remission rates are modest, and the mechanistic basis of non-response is poorly characterised. A common resistance mechanism is thought to emerge when alternative inflammatory cascades compensate for TNF inhibition, but the interactions underlying this rewiring have not been systematically characterised. We applied CytokineLink, our previously developed systems immunology framework, to single-cell RNA sequencing data from CD patients sampled before and after anti-TNF therapy. We reconstructed networks of interacting cytokines across samples stratified by treatment phase, response, and inflammation status, and identified condition-specific cytokine interactions and feedback loops, statistically validated against degree-matched random networks. We clustered the generated networks based on their inflammation, response, and treatment status. The pre-treatment inflamed non-responder network contained the largest set of unique interactions, organised around a connected module driven by IL17C targeting downstream TNF, IL6, IL1B, CXCL1/2/3/8, and CCL20. IL17C was produced by a population of non-ileal enteroendocrine cells, differentially abundant at baseline in non-responders. Gene set variation analysis in an independent cohort confirmed elevated non-responder module activity in colonic tissues of non-responders. Feedback loop analysis revealed that responder networks were characterised by persistent IL10 circuits sustained by macrophage populations and acquired tissue-remodelling interactions after therapy, whereas non-responders lost IL10 feedback loops post-treatment and gained TNF-containing motifs, including circuits signalling through the upstream activator TL1A. Our findings characterise the mechanism of anti-TNF non-response as a cytokine network, in which pre-existing epithelial-driven inflammatory modules and the failure to preserve regulatory feedback sustain TNF-independent inflammation in CD. By characterising cytokine interactions at the systems level, our approach moves beyond single-cytokine models of anti-TNF resistance to provide a mechanistic framework for understanding the biological basis of treatment failure in immune mediated diseases.

17
LiverDCP: A Disease-Cell-Protein Framework for Multi-scale Modeling of Disease Biology

Shi, Z.; Song, Z.; Steveson-Lerner, H.; Dong, B.; Zhao, H.

2026-08-10 systems biology 10.64898/2026.08.07.743628 medRxiv
Top 0.2%
6.6%
Show abstract

Understanding how molecular interactions give rise to disease phenotypes across cellular contexts remains a central challenge in biomedical research. Here, we introduce a Disease-Cell-Protein (DCP) paradigm for modeling multi-scale disease biology, which jointly represents disease states, cellular composition, and protein interaction networks within a unified graph architecture. We instantiate this paradigm in the liver as LiverDCP by integrating a large-scale liver single-cell atlas (LiverHomo) with proteome-wide predicted protein-protein interactions to construct over 280 context-specific interactomes across diverse liver disease and cellular conditions. LiverDCP employs a multi-context representation learning strategy that enables joint training across hundreds of disease-cell environments, capturing shared interaction principles while preserving context-specific variation. LiverDCP incorporates pretrained protein sequence-derived features through a geometry-aware two-phase training scheme that preserves embedding structure while improving predictive performance. The resulting protein embeddings encode context-specific functional states and reveal extensive rewiring of protein roles across diseases. They provide a context-resolved representation of protein function, enabling interpretation of GWAS risk genes and prioritization of therapeutic targets, including recovery of known targets and nomination of candidate repurposed and novel targets for MASH. Overall, this work establishes a generalizable framework for linking molecular interactions to disease phenotypes and enabling mechanistic understanding and target discovery across complex diseases.

18
Single-cell foundation models benefit from cross-modal training: adding proteomics data beats parameter scaling

Burq, M.; Stepec, D.; Kim, C.; Cimermancic, P.

2026-08-19 bioinformatics 10.64898/2026.08.14.744845 medRxiv
Top 0.2%
6.6%
Show abstract

Leading cellular foundation models have been trained on hundreds of millions of single-cell transcriptomes, with progress increasingly driven by larger datasets and model scaling. Here, we asked whether adding a proteomics modality can improve gene-level and cell-level representations beyond scaling RNA-only models. We introduce cross-modal continued pretraining, fine-tuning a published single-cell model (Tahoe-x1) on a large corpus of proteomic profiles. Training a 70M-parameter Tahoe-x1 model for a single epoch on 48843 proteomic samples from 440 diverse mass-spectrometry studies matched or exceeded 1B- and 3B-parameter RNA-only models across most of the original Tahoe-x1 evaluation benchmarks. This shows that with the right training recipe, heterogeneous proteomics data can improve the learned representations of single-cell RNAseq samples, demonstrating strong out-of-distribution generalization. Cross-modal pretraining also improves transfer to a held-out protein perturbation benchmark, where scaling the RNA-only model does not provide comparable benefits. These results demonstrate that careful targeted curation of proteomics data can provide larger benefits than increasing the model size alone and suggest that multimodal pretraining is a promising path toward more informative biological foundation models.

19
Regulatory stochasticity drives opposing phenotypic outcomes in cell-fate decision networks

Hari, K.; Gupta, A.; Shivakumar, L. M.; Kulkarni, P.; Salgia, R.; Jolly, M. K.; Levine, H.

2026-08-26 systems biology 10.64898/2026.08.24.746658 medRxiv
Top 0.2%
6.2%
Show abstract

Gene regulatory network models treat interaction parameters as fixed, although regulatory efficacy fluctuates. We asked how temporal fluctuations in interaction strength reshape phenotype occupancy in cell-fate decision GRN motifs. Across large parameter ensembles, anchored fluctuations largely preserved deterministic occupancies. Additive fluctuations increased occupancy of all-high co-expression states, particularly where high expression was accessible. In contrast, multiplicative fluctuations biased inhibitory interactions toward stronger repression and favored single-high states in a topology-dependent manner. Deterministic controls sampled from noise-induced parameter distributions did not fully reproduce these effects. A Boolean-limit analysis revealed an intrinsic upward bias: loss of repression increased expression regardless of regulator state, whereas stronger repression acted only when the regulator was present. Analyses of epithelial-mesenchymal plasticity and gonadal-fate networks showed increased occupancy of hybrid team-expression states under additive fluctuations. Thus, regulatory noise can reshape the developmental landscape in opposing directions, pushing cell-fate systems toward either progenitor-like or terminally differentiated states.

20
Inferring Protein Variant Impacts Across Contexts

Rasoulzadeh Hosseini, A.; Senguttuvan, V.; van Loggerenberg, W.; Border, R.; Roth, F. P.

2026-08-20 genetics 10.64898/2026.08.18.745369 medRxiv
Top 0.2%
6.1%
Show abstract

Multiplexed assays of variant effects (MAVEs) measure the functional impact of many protein sequence variants in parallel, potentially covering all possible single amino acid substitutions. Unlike current computational variant effect predictors, MAVEs can reveal the effects of variants under different genetic and environmental contexts. However, whereas the space of possible contexts is effectively infinite, contextual MAVE studies are limited by finite experimental budgets. To maximize coverage across contexts, one strategy is to carry out sub-saturation contextual MAVEs and then fill in the gaps via imputation. Here, we categorize and compare different imputation challenges, explore a collection of multi-context imputation solutions, including linear mixed-effects models, random forests, and autoencoders, and provide insight into how best to proceed for a given imputation task. We find that the optimal method depends on the imputation task and how densely the contexts have been measured. More flexible models excel when measurements are plentiful, whereas the simplest models prove most reliable when measurements are sparse. However, the simple source-to-target regression models, although well suited to imputing scores for variants measured in the source context, cannot impute scores for variants that were not measured in either context. This is a major limitation when both maps are sparsely measured. We provide a conceptual framework and an initial evaluation of multi-context imputation methods that can extend the scope of large-scale studies of context-dependent variant effects.